location and appearance
Backdoor Attack in the Physical World
Li, Yiming, Zhai, Tongqing, Jiang, Yong, Li, Zhifeng, Xia, Shu-Tao
Backdoor attack intends to inject hidden backdoor into the deep neural networks (DNNs), such that the prediction of infected models will be maliciously changed if the hidden backdoor is activated by the attacker-defined trigger. Currently, most existing backdoor attacks adopted the setting of static trigger, $i.e.,$ triggers across the training and testing images follow the same appearance and are located in the same area. In this paper, we revisit this attack paradigm by analyzing trigger characteristics. We demonstrate that this attack paradigm is vulnerable when the trigger in testing images is not consistent with the one used for training. As such, those attacks are far less effective in the physical world, where the location and appearance of the trigger in the digitized image may be different from that of the one used for training. Moreover, we also discuss how to alleviate such vulnerability. We hope that this work could inspire more explorations on backdoor properties, to help the design of more advanced backdoor attack and defense methods.
LAVAE: Disentangling Location and Appearance
A BSTRACT We propose a probabilistic generative model for unsupervised learning of structured, interpretable, object-based representations of visual scenes. We use amortized variational inference to train the generative model end-to-end. The learned representations of object location and appearance are fully disentangled, and objects are represented independently of each other in the latent space. Unlike previous approaches that disentangle location and appearance, ours generalizes seam-lessly to scenes with many more objects than encountered in the training regime. We evaluate the proposed model on multi-MNIST and multidSprites data sets. 1 I NTRODUCTION Many hallmarks of human intelligence rely on the capability to perceive the world as a layout of distinct physical objects that endure through time--a skill that infants acquire in early childhood (Spelke, 1990; 2013; Spelke and Kinzler, 2007). Learning compositional, object-based representations of visual scenes, however, is still regarded as an open challenge for artificial systems (Ben-gio et al., 2013; Garnelo and Shanahan, 2019). Recently, there has been a growing interest in unsupervised learning of disentangled representations (Locatello et al., 2018), which should separate the distinct, informative factors of variations in the data, and contain all the information on the data in a compact, interpretable structure (Bengio et al., 2013). This notion is highly relevant in the context of visual scene representation learning, where distinct objects should arguably be represented in a disentangled fashion.
Sequential Attend, Infer, Repeat: Generative Modelling of Moving Objects
Kosiorek, Adam, Kim, Hyunjik, Teh, Yee Whye, Posner, Ingmar
We present Sequential Attend, Infer, Repeat (SQAIR), an interpretable deep generative model for image sequences. It can reliably discover and track objects through the sequence; it can also conditionally generate future frames, thereby simulating expected motion of objects. This is achieved by explicitly encoding object numbers, locations and appearances in the latent variables of the model. SQAIR retains all strengths of its predecessor, Attend, Infer, Repeat (AIR, Eslami et. al. 2016), including unsupervised learning, made possible by inductive biases present in the model structure. We use a moving multi-\textsc{mnist} dataset to show limitations of AIR in detecting overlapping or partially occluded objects, and show how \textsc{sqair} overcomes them by leveraging temporal consistency of objects. Finally, we also apply SQAIR to real-world pedestrian CCTV data, where it learns to reliably detect, track and generate walking pedestrians with no supervision.
Sequential Attend, Infer, Repeat: Generative Modelling of Moving Objects
Kosiorek, Adam, Kim, Hyunjik, Teh, Yee Whye, Posner, Ingmar
We present Sequential Attend, Infer, Repeat (SQAIR), an interpretable deep generative model for image sequences. It can reliably discover and track objects through the sequence; it can also conditionally generate future frames, thereby simulating expected motion of objects. This is achieved by explicitly encoding object numbers, locations and appearances in the latent variables of the model. SQAIR retains all strengths of its predecessor, Attend, Infer, Repeat (AIR, Eslami et. al. 2016), including unsupervised learning, made possible by inductive biases present in the model structure. We use a moving multi-\textsc{mnist} dataset to show limitations of AIR in detecting overlapping or partially occluded objects, and show how \textsc{sqair} overcomes them by leveraging temporal consistency of objects. Finally, we also apply SQAIR to real-world pedestrian CCTV data, where it learns to reliably detect, track and generate walking pedestrians with no supervision.